Papers with synthesis quality

7 papers
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding (2025.acl-demo)

Copied to clipboard

Challenge: Experimental evaluations show RT-VC delivers a 13.3% reduction in latency . voice conversion modifies speech to match the timbre of a target speaker while preserving content information.
Approach: They propose a zero-shot real-time voice conversion system that leverages an articulatory feature space to naturally disentangle content and speaker characteristics.
Outcome: The proposed system achieves a CPU latency of 61.4 ms, representing a 13.3% reduction in latency.
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models fail to generate singing voices rich in stylistic nuances for unseen singers due to multifaceted nature of singing styles.
Approach: They propose a zero-shot SVS model for style transfer across cross-lingual speech and singing styles and multi-level style control.
Outcome: Experimental results show that TCSinger outperforms baseline models in synthesis quality, singer similarity, and style controllability.
Empowering Diffusion Models on the Embedding Space for Text Generation (2024.naacl-long)

Copied to clipboard

Challenge: Recent work adapts diffusion models to textual data by diffusing on the embedding space.
Approach: They propose an embedding diffusion model based on Transformer to solve the problem of embeddable space and denoising model.
Outcome: The proposed model is more efficient than previous methods on seminal text generation tasks and is superior to existing models.
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature (2026.eacl-long)

Copied to clipboard

Challenge: Existing retrieval-augmented generation methods overlook citation graph structure, adapt poorly to complex queries, and yield fragmented, hard-to-verify syntheses.
Approach: They propose a retrieval-augmented generation framework that addresses these gaps by combining adaptive retrieval and symbolic reasoning.
Outcome: Extensive experiments show that SciRAG outperforms prior systems in factual accuracy and synthesis quality.
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)

Copied to clipboard

Challenge: Text-to-speech (TTS) models have been developed to generate high-quality speech.
Approach: They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively.
Outcome: The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing.
MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Text-to-Audio Generation (TTA) systems suffer from slow inference speed, authors report . authors demonstrate that MeanAudia achieves state-of-the-art performance in single-step audio generation .
Approach: They propose a text-to-audio generator capable of rendering realistic sound with only one function evaluation.
Outcome: The proposed system achieves state-of-the-art performance in single-step audio generation.
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis (2025.findings-acl)

Copied to clipboard

Challenge: Existing zero-shot singing voice synthesis models depend on phoneme and note boundary annotations, limiting their robustness and producing poor transitions between phonemes and notes.
Approach: They propose a multi-task multilingual zero-shot SVS model with style transfer and style control based on various prompts.
Outcome: Experimental results show that TCSinger 2 outperforms baseline models in subjective and objective metrics across multiple related tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations